Skip to content

UN-4223 [FIX] Make streamed Gemini calls honour the adapter timeout - #2313

Open
johnyrahul wants to merge 2 commits into
mainfrom
UN-4223-vertex-stream-timeout
Open

johnyrahul wants to merge 2 commits into
mainfrom
UN-4223-vertex-stream-timeout

Conversation

@johnyrahul

Copy link
Copy Markdown
Contributor

What

  • Streamed Gemini completions (gemini/* via the Gemini adapter, vertex_ai/*gemini* via the Vertex AI adapter) now apply the adapter's timeout to the HTTP request. Before this change the request used LiteLLM's 6000 s default.
  • Both streaming paths are covered: LLM.complete() (which streams by default) and LLM.stream_complete().

Why

UN-4223: on 2026-10-01, 10-02 and 10-03, single vertex_ai/gemini-3.1-flash-lite calls in unstract-worker-pg-executor hung for 2+ hours, and pod liveness eventually took whole executor pods offline (P0 on the US Checkly check). The SDK sets timeout=600 on these calls, but it never took effect.

The root cause is in LiteLLM 1.96.2 (vertex_and_google_ai_studio_gemini.py):

  • On the sync streaming path, make_sync_call is bound without timeout and posts on litellm.module_level_client.
  • That client is built with litellm.request_timeout, which defaults to 6000 s.
  • The non-streaming branch of the same function does honour timeout. Streaming has been on by default since we added it, so in practice every Gemini call was capped at 100 minutes per read instead of 10.

Anthropic's handler forwards timeout per request, so it is unaffected. I checked this in the same LiteLLM version.

This is not the whole fix for UN-4223. Prod logs show the hung Gemini calls never timed out even at 6000 s. No Retry, timeout or error line was logged over 2+ hours, while memory climbed to around 9–10 GB. That points to a stream that keeps sending data, which a per-read timeout cannot catch. A total wall-clock deadline on streamed completions follows in a separate PR. This PR closes a real gap that was hiding behind it: a genuinely silent Gemini stream should fail in 10 minutes, not 100.

How

  • _with_gemini_stream_timeout() adds client=HTTPHandler(timeout=httpx.Timeout(<adapter timeout>)) to the streaming call's kwargs when:
    • the model is a Gemini model by LiteLLM's own routing (gemini/ prefix, or vertex_ai/ with gemini in the name), and
    • a positive timeout is set, and
    • the caller has not already passed a client.
  • The handler only accepts a client as an HTTPHandler (gemini_client=client if isinstance(client, HTTPHandler)), so passing a client is the only way to get the timeout through.
  • One client is cached per timeout value, so connections stay pooled like module_level_client.
  • The caller's kwargs are not mutated, and non-streaming paths are untouched.

Can this PR break any existing features. If yes, please list possible items. If no, please explain why. (PS: Admins do not merge the PR without this section filled)

  • Behaviour change, intended: a streamed Gemini request that goes silent for longer than the adapter's timeout (default 600 s; the Vertex form exposes no timeout field) now fails and retries instead of waiting up to 6000 s. A healthy stream sends chunks continuously and the read timeout resets on each one, so long generations are unaffected.
  • Vertex partner models (e.g. vertex_ai/claude-*) take LiteLLM's partner route, which already forwards timeout. They don't match the Gemini check and are pinned by a test.
  • Every other provider is untouched (pinned for Anthropic and OpenAI).

Database Migrations

  • None.

Env Config

  • None.

Relevant Docs

  • None.

Related Issues or PRs

  • UN-4223 (prod outages), UN-4179 (same mechanism on staging, Anthropic).
  • Follow-up: a total wall-clock deadline for streamed completions (separate PR).

Dependencies Versions

  • None. LiteLLM stays at 1.96.2.

Notes on Testing

  • New tests/test_gemini_stream_timeout.py with 11 cases.
  • An end-to-end test through LiteLLM stubs the HTTP transport and asserts the outgoing request's read timeout equals the adapter's. Mutation-verified: with the fix removed, the same test sees 6000.0, reproducing the prod bug, and the Vertex adapter test fails on the missing client.
  • Unit cases: which models get a client and which don't (Anthropic, OpenAI, Vertex partner), no client when there's no usable timeout, a caller-supplied client is kept, and the client is shared per timeout.
  • Full sdk1 suite: 731 passed. ruff 0.3.4 (the pre-commit pin) is clean.

Screenshots

Checklist

I have read and understood the Contribution Guidelines.

🤖 Generated with Claude Code

LiteLLM's Gemini handler (gemini/* and vertex_ai/*gemini*) drops the
per-call timeout on the sync streaming path and streams on
litellm.module_level_client, whose deadline is litellm.request_timeout
(6000 s by default). The adapter's timeout (600 s) never reached the
request, so a stalled Gemini stream could sit for 100 minutes per attempt.

Pass a shared HTTPHandler built with the adapter's timeout for Gemini
models on both streaming paths (complete and stream_complete). Other
providers already forward timeout to the request and are untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@greptile-apps

greptile-apps Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

via Greptile

RetriggerConfidence Score: 5/5

[Medium risk] Adds timeout handling for streamed Gemini API calls.

The PR appears safe to merge; no new actionable issue was established.

Summary

The PR passes adapter timeouts to streamed Gemini and Vertex Gemini HTTP requests and bounds the per-timeout client cache.

  • New request-level tests cover Gemini completion, Vertex completion, and Gemini streaming completion.

Reviews (2) · Last reviewed commit: "UN-4223 [FIX] Bound the Gemini stream cl..."

Comment thread unstract/sdk1/src/unstract/sdk1/llm.py Outdated
Comment thread unstract/sdk1/tests/test_gemini_stream_timeout.py Outdated
Cap the per-timeout client cache at 8 entries so a worker that sees many
distinct adapter timeouts does not hold a connection pool for each.

Replace the mock-only Vertex test with request-level tests: the HTTP
request's read timeout is now asserted for the Vertex adapter (auth
stubbed) and for stream_complete(), as well as the Gemini adapter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@sonarqubecloud

sonarqubecloud Bot commented Oct 5, 2026

Copy link
Copy Markdown

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Unstract test results

Per-group results

Status Group Tier Passed Failed Errors Skipped Duration (s)
✅ e2e-api-deployment e2e 3 0 0 0 8.5
✅ e2e-coowners e2e 1 0 0 0 1.4
✅ e2e-etl e2e 1 0 0 0 16.7
✅ e2e-login e2e 2 0 0 0 1.2
✅ e2e-prompt-studio e2e 1 0 0 0 11.6
✅ e2e-smoke e2e 2 0 0 0 1.1
✅ e2e-workflow e2e 1 0 0 0 14.3
✅ frontend unit 620 0 0 0 14.1
✅ integration-backend integration 603 0 0 26 57.9
✅ integration-connectors integration 1 0 0 7 8.5
✅ integration-workers integration 164 0 0 1 39.9
❌ ui e2e 0 1 0 0 0.0
✅ unit-backend unit 1422 0 0 1 36.9
✅ unit-connectors unit 72 0 0 0 8.8
✅ unit-core unit 276 0 0 0 2.5
✅ unit-platform-service unit 15 0 0 0 2.2
✅ unit-rig unit 120 0 0 0 4.1
✅ unit-runner unit 10 0 0 0 3.3
✅ unit-sdk1 unit 730 0 0 0 29.1
✅ unit-workers unit 1383 0 0 1 111.9
TOTAL 5427 1 0 36 374.2

Critical paths

⚠️ Critical paths not yet covered

  • workflow-execution-fan-out — Multi-file workflow execution fans out to file-processing workers and rejoins. (declared coverage: no groups declared)
✅ Covered critical paths
  • auth-login — covered by e2e-login
  • adapter-register-llm — covered by integration-backend
  • workflow-author — covered by integration-backend
  • co-owner-manage — covered by integration-backend, e2e-coowners
  • workflow-create-execute — covered by e2e-workflow
  • api-deployment-provision — covered by integration-backend
  • api-deployment-auth — covered by integration-backend
  • api-deployment-run — covered by e2e-api-deployment
  • mcp-server-auth — covered by integration-backend
  • mcp-platform-auth — covered by integration-backend
  • platform-key-whoami — covered by integration-backend
  • prompt-studio-author — covered by integration-backend
  • prompt-studio-fetch-response — covered by e2e-prompt-studio
  • connector-register-test — covered by integration-backend
  • pipeline-etl-execute — covered by e2e-etl
  • usage-aggregate-read — covered by integration-backend
  • usage-token-tracking — covered by e2e-api-deployment
  • callback-result-delivery — covered by e2e-api-deployment

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant